Papers with Mozilla Common Voice
Samrómur: Crowd-sourcing large amounts of data (2022.lrec-1)
Copied to clipboard
| Challenge: | Samrómur is the largest prompted speech collection effort for Icelandic so far and verification is as monumental as the collection itself. |
| Approach: | They propose to collect large and diverse corpus for automatic speech recognition and similar tools using crowd-sourced donations. |
| Outcome: | The collected utterances are based on the Mozilla Common Voice platform and are available for free on the Samrómur collection platform. |
Accented Speech Recognition With Accent-specific Codebooks (2023.emnlp-main)
Copied to clipboard
| Challenge: | Degradation in performance across underrepresented accents is a severe deterrent to inclusive adoption of ASR. |
| Approach: | They propose an approach to adapt speech accents to unseen accents by using cross-attention with a trainable set of codebooks. |
| Outcome: | The proposed approach yields significant performance gains on the seen English accents and unseen accents on the Mozilla Common Voice dataset. |
Artie Bias Corpus: An Open Dataset for Detecting Demographic Bias in Speech Applications (2020.lrec-1)
Copied to clipboard
| Challenge: | A speech technology exhibits demographic bias when performance is worse for one demographic group relative to another. |
| Approach: | They create an English dataset of expert-validated audio, transcript> pairs with demographic tags for age, gender, accent and open software which may be used to detect demographic bias in Automatic Speech Recognition systems. |
| Outcome: | The Artie Bias Corpus is a curated subset of the Mozilla Common Voice corpus, which is released under a Creative Commons CC0 license . |